fix(ci): the invisible-character gate never matched anything - #85
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🔇 Additional comments (3)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe workflow updates its invisible-character pattern to use Unicode code-point escapes. It also passes ChangesInvisible-character gate
Estimated code review effort: 1 (Trivial) | ~5 minutes Merge Risk: ⚪ Minimal · up to The workflow change is localized to correcting invisible-character matching and presents no actionable merge-blocking risk at the current head; it is merge-ready after normal checks and review. Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Description checkExplanation The description provides a clear summary, root cause, implementation details, and verification evidence. It does not reproduce the repository checklist or use all template headings, but it is otherwise substantially complete. Full details: Linked Issues checkExplanation The change fixes code-point escapes, adds C0 control detection, and uses grep -a. It does not implement the required separate leading-BOM check or update the compiled linter and configuration for consistent C0 detection. Resolution Add the separate byte-wise leading-BOM check. Update stdlib/ByteDetector.affine and config.ncl with the matching is_c0_control/1 logic. Add or run verification for all listed invisible characters, NUL-bearing files, clean files, permitted whitespace, and the corrupted workflow. Apply the correction to all required gate copies if this issue covers the wider estate. Full details: Docstring CoverageExplanation No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
The PR successfully identifies that the previous CI gate was ineffective, but the proposed fix using \x{...} escapes in grep -P introduces new risks. These escapes often match single bytes or require specific UTF-8 modes to function correctly, which may lead to the gate failing silently again. It is recommended to revert to explicit UTF-8 byte sequences for maximum robustness.
Furthermore, the scanning command can be optimized for performance and better debuggability by using find's batching capabilities and enabling error visibility. A major gap exists in the testing strategy; without adding sample files containing the target invisible characters to the repository, there is no way to ensure these patterns work now or in the future.
About this PR
- The PR lacks automated regression tests. While manual verification was performed, no sample files containing these invisible characters (NBSP, BOM, C0 controls, etc.) were added to the repository to verify the gate or prevent future regressions.
Test suggestions
- Missing recommended test scenario: Detect Non-breaking Space (U+00A0) using the new codepoint escape
- Missing recommended test scenario: Detect C0 control characters (e.g., backspace \x08) in source files
- Missing recommended test scenario: Detect Byte Order Mark (BOM) (U+FEFF) within or at the start of files
- Missing recommended test scenario: Detect Zero-width characters (U+200B, U+200C, etc.) and Bidi overrides (U+202A-U+202E)
- Missing recommended test scenario: Verify scanning of files containing NUL bytes (\x00) using the '-a' flag to avoid binary-skip behavior
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Missing recommended test scenario: Detect Non-breaking Space (U+00A0) using the new codepoint escape
2. Missing recommended test scenario: Detect C0 control characters (e.g., backspace \x08) in source files
3. Missing recommended test scenario: Detect Byte Order Mark (BOM) (U+FEFF) within or at the start of files
4. Missing recommended test scenario: Detect Zero-width characters (U+200B, U+200C, etc.) and Bidi overrides (U+202A-U+202E)
5. Missing recommended test scenario: Verify scanning of files containing NUL bytes (\x00) using the '-a' flag to avoid binary-skip behavior
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| # non-breaking spaces, null bytes, and other invisible Unicode in source files. | ||
| set +e | ||
| PATTERNS='\xc2\xa0|\xe2\x80\x8b|\xe2\x80\x8c|\xe2\x80\x8d|\xef\xbb\xbf|\xc2\xad|\xe2\x80\x8e|\xe2\x80\x8f|\xe2\x80\xaa|\xe2\x80\xab|\xe2\x80\xac|\xe2\x80\xad|\xe2\x80\xae|\x00' | ||
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' |
There was a problem hiding this comment.
🔴 HIGH RISK
The \x{...} syntax is likely to break the detection logic or cause silenced errors in this environment. Use UTF-8 byte sequences instead to ensure the gate correctly identifies invisible characters. Additionally, consider including directional overrides (U+2066-U+2069) often used in 'Trojan Source' attacks.
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}' | |
| PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\xc2\xa0|\xc2\xad|\xe2\x80\x8b|\xe2\x80\x8c|\xe2\x80\x8d|\xe2\x80\x8e|\xe2\x80\x8f|\xe2\x80\xaa|\xe2\x80\xab|\xe2\x80\xac|\xe2\x80\xad|\xe2\x80\xae|\xe2\x81\xa0|\xef\xbb\xbf' |
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
⚪ LOW RISK
Suggestion: The scanning command can be optimized for performance and clarity. Using + instead of \; allows grep to process multiple files per invocation, and the -r flag is unnecessary when find provides the file list. Also, removing 2>/dev/null ensures that any issues with the regex engine are visible in the logs.
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt |



Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.